Papers with visual-textual alignment

2 papers
Granularity Matters: Pathological Graph-driven Cross-modal Alignment for Brain CT Report Generation (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods for automatic Brain CT reports are limited by coarse-grained supervision and coupled cross-modal alignment.
Approach: They propose a pathological Graph-driven cross-modal alignment model that learns fine-grained visual cues and aligns them with textual words.
Outcome: The proposed model can improve the automatic generation of Brain CT reports and contribute to improved cranial disease diagnosis.
VLA-Mark: A cross modal watermark for large vision-language alignment models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing text watermarking methods disrupt visual-textual alignment, leaving semantic-critical concepts vulnerable.
Approach: They propose a vision-aligned framework that embeds detectable watermarks into outputs . they combine localized patch affinity, global semantic coherence, contextual attention patterns .
Outcome: The proposed framework shows lower PPL and higher BLEU than conventional methods with near-perfect detection (98.8% AUC).

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations